Back

NAR Genomics and Bioinformatics

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match NAR Genomics and Bioinformatics's content profile, based on 242 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit.

1
Targeted sequencing of expanded tandem repeats: Identifying interruptions and errors

Heuchan, A. E. R.; Peter Durairaj, R. R.; Aston, A. N.; Randall, E. L.; Heraty, L.; Smith, C.; Taylor, A. S.; Massey, T. H.; Tome, S. H.; Dion, V.

2026-07-26 bioinformatics 10.64898/2026.07.23.739298 medRxiv
Top 0.1%
22.2%
Show abstract

Expanded repeat disorders remain without a disease modifying treatment. Some of the most important modifiers of disease severity include inherited repeat size, their somatic mosaicism, and whether they contain interruptions. Targeted sequencing approaches are becoming the gold standard on how these parameters are measured despite repeats being notoriously challenging to sequence. Here we developed CHARLIE (Comprehensive High-throughput Analysis of Repeat Length, Interruptions and Expansions) to identify and position interruptions within expanded repeats. We used samples from Huntingtons disease and myotonic dystrophy type 1 sequenced with Illumina MiSeq and PacBios Single-Molecule Real-Time Sequencing. We validated the pipeline against previous methods for identifying germline-inherited interruptions. One challenge that can potentially mask the identification of true interruptions and sequence variation is the accuracy of sequencing platforms, which, when applied to expanded repeats, is largely unknown. Our results suggest that each sequencing platform produces a distinct error profile and we show that PCR-free library preparation for SMRT sequencing improves sequencing accuracy. CHARLIE, therefore, provides a method that can help differentiate between trivial sequencing errors and those interruptions that modify disease presentation, which can be inherited or somatic in origin.

2
FERRET: Framework to Evaluate Robustness in Regulatory Networks Using Heterogeneous Cell Types

Eicher, T. D.; Quackenbush, J.

2026-07-25 bioinformatics 10.64898/2026.07.21.739795 medRxiv
Top 0.1%
21.9%
Show abstract

Techniques for evaluating gene regulatory network (GRN) inference methods typically focus on recovering small ground-truth networks or on benchmarking against simulated data. However, both approaches have important limitations and fail to capture the biological variability present in real datasets. FERRET is a framework for benchmarking single-cell GRN inference methods based on a simple biological assumption: independent estimates of the regulatory network from the same cellular state should resemble one another more closely than estimates from distinct cellular states. Rather than relying on incomplete or simulated ground truth, FERRET quantifies network robustness using two complementary metrics: Robustness Area Under the Curve (RAUC), an AUC-like measure of within-cell-type network similarity relative to between-cell-type similarity, and Monotonicity, which assesses the consistency of network similarity across edge-weight cutoffs. FERRET also supports biological validation through pathway enrichment analysis. We validate FERRET using experimentally derived ChIP-seq networks from B lymphocytes and fibroblasts as positive controls and randomly generated networks as negative controls, showing that biologically related networks receive high robustness scores whereas randomly generated networks receive scores consistent with chance. Finally, we apply FERRET to multiple GRN inference methods on real single-cell RNA-sequencing datasets to identify methods that produce the most robust, biologically informative regulatory networks. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=112 SRC="FIGDIR/small/739795v1_ufig1.gif" ALT="Figure 1"> View larger version (43K): org.highwire.dtl.DTLVardef@da7e4eorg.highwire.dtl.DTLVardef@9a5425org.highwire.dtl.DTLVardef@a6e60org.highwire.dtl.DTLVardef@d466cd_HPS_FORMAT_FIGEXP M_FIG C_FIG

3
Method Choice, Not Biology, Determines In Silico Perturbation Results: A Systematic Evaluation of Eight Methods Across Four Datasets

Wenjie, G.; Wu, S.; Hu, G.; Yang, Z.; Wang, Z.; Cai, J.; Mao, J.

2026-08-19 bioinformatics 10.64898/2026.08.11.744106 medRxiv
Top 0.1%
21.5%
Show abstract

Most in silico perturbation methods for single-cell transcriptomics have been validated only on individual datasets, leaving their reliability and generalizability unknown. Through systematic cross-method, cross-dataset benchmarking of eight methods spanning six mathematical frameworks across four datasets, we find that six of eight methods--including widely used VAE-based and tensor decomposition approaches--fail to produce detectable transcription factor (TF)-to-pathway signals. Only CellOracle and DDIM consistently detected TF-to-glycolysis directional regulation. Cross-pathway analysis in PBMC monocytes revealed biologically coherent TF-pathway associations beyond glycolysis (SPI1[->]glycolysis 4.4x enrichment, FOS[->]AP-1 targets 4.4x), with SOX9 serving as a biological specificity control (no pathway enrichment). Method choice alone could reverse biological conclusions: DDIM and scTenifoldKnk rankings were significantly anti-correlated ({rho}=-0.811, p=0.027). CRISPRi Perturb-seq validation in K562 cells confirmed TF knockdown suppresses glycolysis gene expression (JUN {delta}=-1.72, CEBPB {delta}=-1.59, SPI1 {delta}=-1.57, FOS {delta}=-0.70), but CellOracle-predicted perturbation directions did not match experimental directions (40.9% agreement, not different from chance), revealing a fundamental gap between steady-state correlation and causal perturbation. Diagnostic analyses using VAE latent space profiling, correlation distribution comparison, and gene-gene graph analysis identified distinct failure modes in unsuccessful methods: VAE latent space competition (STAT3 signal-to-noise 0.44 vs. SPI1 4.25), correlation noise (TF-glycolysis |r|=0.038 indistinguishable from background |r|=0.047), and graph non-specificity (0.84x enrichment). A controlled ablation experiment showed that adding a GRN prior to DDIM did not improve target recall (delta=0 for all TFs), confirming that performance differences are multi-factorial. These findings establish preliminary guidance for method selection, including cross-pathway validation, direction-aware benchmarking, and minimum data requirements ([≥]500 cells, [≥]1,000 HVGs).

4
Integrating suboptimal secondary structures, AI-assisted genomic synteny, and evolutionary conservation to identify bacterial ncRNA homologs beyond sequence similarity

Panek, J.

2026-07-16 bioinformatics 10.64898/2026.07.10.737749 medRxiv
Top 0.1%
18.8%
Show abstract

A bioinformatic approach for genome-wide identification of homologs of bacterial non-coding RNAs (ncRNAs) integrating structural similarity, genomic synteny, and evolutionary conservation is presented. The structural similarity is detected using an algorithm for genome-wide identification of loci in genomic intergenic regions (IGRs) containing sequences capable of adopting secondary structures similar to that of the query ncRNA. The algorithm scans IGR sequences using a sliding window with a predefined step. For each window, suboptimal secondary structures are predicted and compared with the template structure to compute structural similarity scores. These scores are evaluated statistically on a genome-wide scale to infer homology of the RNAs represented by the predicted structures. Loci encoding statistically significant structures are further filtered using genomic synteny of the query ncRNAs inferred from genomic annotations. ChatGPT was used to assist in identifying literature-supported biological relationships between genes with distinct functional annotations. Syntenic loci with the structures are then examined for homologs in related species, as evolutionary conservation among related species is a common feature of ncRNAs Using this approach, we predicted novel homologs of the spot42 RNA-encoding spf gene in Glaciecola and Pseudoalteromonas genomes, and ms1 RNA genes in Frankia and Bifidobacterium genomes, where previous homology searches had failed. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=44 SRC="FIGDIR/small/737749v1_ufig1.gif" ALT="Figure 1"> View larger version (13K): org.highwire.dtl.DTLVardef@dd18baorg.highwire.dtl.DTLVardef@18261eforg.highwire.dtl.DTLVardef@eb95c3org.highwire.dtl.DTLVardef@b535ad_HPS_FORMAT_FIGEXP M_FIG C_FIG

5
SpliSync: Genomic language model-driven splice site correction of long RNA sequencing reads

Lui, W. W.; Florea, L. D.

2026-07-06 bioinformatics 10.64898/2026.07.04.736518 medRxiv
Top 0.1%
15.1%
Show abstract

Long RNA sequencing reads are rapidly replacing short reads in transcriptomic analyses, enabling full-length transcript sequencing and better identification of isoforms, alternative splicing events, and other transcript variants. However, their higher sequencing error rates can cause misalignments, especially at splice junctions, reducing the accuracy of transcript reconstruction and analysis. We developed SpliSync, a genomic language model-driven method for splice site correction that integrates a pre-trained genomic sequence model (HyenaDNA), alignment data, and a U-net architecture to predict splice sites at nucleotide resolution. SpliSync substantially improved the precision of RNA long-read alignments by 27%-194% across diverse datasets and consistently outperformed competing tools. As a preprocessing step, it increased alternative splicing detection accuracy by 26%-330%. In contrast, its benefit for transcript reconstruction was limited, likely due to the tools' built-in correction mechanisms. The code, developed in Python using the PyTorch package, is freely available at https://github.com/splicebox/SpliSync.

6
Hyper3D-lite: count-preserving representation auditing for long-read multi-contact genome data

Zhang, Y.

2026-06-11 bioinformatics 10.64898/2026.06.09.730992 medRxiv
Top 0.1%
14.7%
Show abstract

Long-read and single-molecule sequencing technologies are rapidly increasing molecule-level data, with platforms such as Oxford Nanopore, PacBio HiFi, and Roche sequencing-by-expansion advancing at different technology readiness levels. In the specific context of Pore-C and HiPore-C multi-contact chromatin-conformation assays, long-read multi-contact 3D genome assays preserve molecule-level contact context, but common downstream pairwise projections can expand one multi-contact molecule into many pair records. This creates a representation problem: apparent contact evidence can increase through the counting frame before biological interpretation begins. Hyper3D-lite addresses this problem as a representation-first audit tool for read-to-fragment-style long-read multi-contact inputs. It compares all-pair projection with CPB, a count-preserving statistical accounting reference point, and separates broad software outputs from conservative higher-order candidate calls.

7
ORBIT: Annotation-Aware Empirical Enrichment and Semantic Reranking for Interpretable Functional-Class Recovery

Kidder, B. L.

2026-07-07 bioinformatics 10.64898/2026.07.01.735870 medRxiv
Top 0.1%
14.6%
Show abstract

Gene-set interpretation workflows are widely used to summarize transcriptomic and proteomic experiments, yet standard enrichment tools often return long, redundant result tables that require substantial manual consolidation. We developed ORBIT (Ontology-Ranked Biological Interpretation Tool), an annotation-aware interpretation workflow that combines empirical enrichment, semantic reranking, and redundancy-aware representative-term selection to prioritize interpretable functional summaries from gene sets. We evaluated ORBIT on a curated tiered benchmark of human functional-class gene sets spanning clean reference sets, size-ladder variants, and mixed-difficulty cases. On the 45-set core benchmark, ORBIT semantic achieved higher expected-class recovery than Enrichr and PANTHER Gene Ontology molecular-function baselines, with a mean reciprocal rank of 0.916 and top-1 recovery of 0.889. Bootstrap confidence intervals and paired permutation testing supported the robustness of this advantage, and supplemental analyses extended the comparison to g:Profiler. In a GPCR mixed-function case study, ORBIT compressed redundant enriched terms into semantic representative neighborhoods, illustrating how long enrichment outputs can be converted into reviewable biological summaries. We then used ORBIT to interpret immune-cell identity, interferon-response biology, and breast-cancer subtype programs. ORBIT linked PBMC3K markers to cytotoxic, antigen-presentation, and innate-immune cell states; prioritized antiviral, cytokine-response, RNA-binding, and secreted-factor biology after IFNB stimulation; and separated TCGA-BRCA basal-like proliferative chromosome/cell-cycle programs from luminal transporter and receptor-associated biology while retaining gene-level support.

8
How do noise and precision contraints jointly determine the minimum sketch size for fixed-size MinHash Jaccard estimation?

Ebou, A. E. T.; KOUA, D. K.

2026-08-20 bioinformatics 10.64898/2026.08.15.744971 medRxiv
Top 0.1%
14.5%
Show abstract

MinHash-based genome comparison is governed by two statistical constraints: a noise floor on k-mer size, controlling chance collisions between unrelated sequences, and a precision floor on sketch size, controlling uncertainty in the estimated Jaccard similarity. These constraints are best characterized for fixed-size bottom-sketch MinHash, the classical construction underlying tools such as Mash and still the standard basis for comparing genomes of similar size. Nevertheless, while, the noise floor has an established closed-form solution, sketch size is commonly selected using fixed defaults or heuristic choices independent of genome length and target precision. Moreover, recent work on scaled (FracMinHash) sketching has highlighted the difficulty of obtaining an analogous closed-form confidence interval for the Jaccard similarity. Here, we show that under the fixed-size bottom-sketch MinHash model, the shared-hash count follows a binomial model, which lets classical proportion-estimation theory be applied directly. Combining this with Fofanov's k-selection criterion and the exact identity-Jaccard relationship yields a single closed-form design rule, s* (L, {rho}* , ), that unifies the noise and precision floors into one operating envelope. The resulting design rule was validated against exact Clopper-Pearson interval inversion, idealized binomial sampling, realistic overlapping k-mer simulation and 54 bacterial genome pairs spanning species-to-order taxonomic levels. In realistic simulations, empirical coverage remained within 1-3 percentage points of the idealized reference in 14 of 15 tested cells. In real genomes, the classical identity-Jaccard relationship showed increasing positive bias with taxonomic divergence, from a median of -1.4% within species to +27% at genus and +77% at family level. We further show that s* (L) is not smooth in genome length but a discontinuous, previously unreported staircase caused by the ceiling function used for k-mer selection, with practical consequences concentrated at specific genome-size boundaries.

9
FetchR: an intuitive end-to-end solution for local RNA-Seq analyses

Fetch, D. R.; Soshnev, A. A.

2026-08-04 genomics 10.64898/2026.07.30.741593 medRxiv
Top 0.1%
14.5%
Show abstract

RNA-Seq, analyses of RNA abundance by next-generation sequencing, has become a near-universal tool in modern biology. Availability of streamlined protocols and kits, straightforward ability to multiplex hundreds of samples, low cost of short-read sequencing, and well-established analytical pipelines make RNA-Seq a method of choice when even a few genes need to be analyzed in parallel. While many tools have been developed for quality control, mapping, and visualization of RNA-Seq data, managing all these individually still requires substantial familiarity with shell scripting and R, and remains a bottleneck for laboratories with limited computational background. We assembled FetchR, an intuitive pipeline with built-in, clear explanations of features and outputs, for local analyses of RNA-Seq data from either own .fastq files or data imported from Sequence Read Archive via the ENA Portal API. The pipeline operates in Windows Subsystem for Linux (WSL) and is installed via a single script that handles all individual tools, as well as their dependencies and updates, including the reference genome annotation(s), and system requirements. The outputs include standard quality control checks, data visualization, read summation, differential gene expression analyses, visualization, and exploratory analyses using Gene Ontology and Gene Set Enrichment Analyses, as well as detailed logs of every step for subsequent reproducible reporting.

10
Systematic assessment of sequencing depth requirements for Hi-C-derived metrics

Granovsky, A.; Polovnikov, K. E.

2026-08-01 bioinformatics 10.64898/2026.07.28.741297 medRxiv
Top 0.1%
13.6%
Show abstract

BackgroundHi-C experiments produce genome-wide chromatin contact maps from which structural features can be quantified across multiple genomic scales, ranging from megabase-scale compartments to kilobase-scale boundaries and chromatin loops. Although high-resolution analyses commonly rely on hundreds of millions of sequenced read pairs, the minimum sequencing depth required for different classes of Hi-C-derived features has not been systematically established. We therefore sought to determine these depth requirements systematically. ResultsWe performed progressive random subsampling of twelve Hi-C libraries representing multiple cell types and experimental protocols. Loop-density and loop-size inference from the P(s) log-derivative was evaluated across all libraries, whereas compartment, insulation, and boundary analyses were performed on a subset of nine libraries with comparable full-depth coverage. Rather than defining sufficient depth as an arbitrary fraction of the full-depth value, we introduced a biologically motivated criterion based on the variability between independent biological replicates. The required sequencing depth for each metric was defined as the point at which subsampling scatter first reached the variability observed between independent biological replicates, representing the accuracy that additional sequencing cannot improve upon. Using this criterion, loop-density estimation reached its biological floor at approximately 10 million read pairs, loop-size estimation at approximately 20 million, insulation scores and boundary detection at approximately 30 million, and compartment eigenvectors at 100-kilobase resolution at approximately 60 million read pairs. ConclusionsDifferent Hi-C-derived metrics require substantially different sequencing depths to achieve biologically meaningful accuracy. For all metrics, depth-induced variability fell below biological replicate variability well before full sequencing depth was reached. These thresholds provide practical guidance for experimental design and sequencing budget allocation, suggesting that, in many studies, increasing the number of biological replicates is likely to improve reproducibility more effectively than sequencing individual libraries to greater depth.

11
DDTRN: Predicting Bacterial Transcriptional Regulatory Networks Based on Gene Sequences using Dual Descriptor

Nie, P.; Ma, B.-G.

2026-07-01 bioinformatics 10.64898/2026.06.30.735580 medRxiv
Top 0.1%
13.0%
Show abstract

Accurate computational reconstruction of bacterial transcriptional regulatory network (TRN) from sequence information alone remains a fundamental challenge in systems biology, particularly for non-model organisms lacking extensive transcriptomic data. We present DDTRN, a sequence-driven framework that formulates TRN inference as a binary classification task over concatenated regulator-target gene sequence pairs and employs a Dual Descriptor (DD) model to predict regulatory interactions. The DD architecture represents a sequence into two learnable components: Composition Weight Map (CWM) and Position Weight Function (PWF). We comprehensively evaluate DDTRN against six conventional machine learning baselines across eight benchmark bacterial datasets, including E. coli (DREAM5, RegulonDB), B. subtilis, S. enterica, C. glutamicum, M. tuberculosis, P. aeruginosa, and S. coelicolor. DDTRN achieves superior overall performance, attaining average AUROC and AUPR scores of 0.869 and 0.868, respectively, with particularly pronounced advantages at lower descriptor ranks where positional weighting compensates for limited sequence context. Systematic sensitivity analyses of rank, embedding dimension, and basis function count reveal stable optimal operating regimes, while subsampling experiments demonstrate strong robustness even with limited training data. Interpretability analyses show that PWF learns distinct periodic contributions across different rank granularities and that CWM preferentially weights meaningful k-mers. A case study on E. coli dataset further illustrates that DDTRN identifies method-specific candidate targets complementary to those proposed by conventional approaches. By operating solely on genomic sequence, DDTRN provides a scalable, interpretable, and data-efficient framework for bacterial TRN inference in species where expression data are scarce, and it establishes a foundation for future multimodal integration with condition-specific regulatory information.

12
Scanning transcriptomes for nonlinear, domain-level similarities using hmSEEKR

Li, S.; Sprague, D. A.; Eberhard, Q. E.; Boyson, S. P.; Laederach, A.; Calabrese, J. M.

2026-07-08 bioinformatics 10.64898/2026.07.03.736302 medRxiv
Top 0.1%
12.8%
Show abstract

Long noncoding RNAs (lncRNAs) play roles in gene regulation across kingdoms of life. However, lncRNAs with related functions often lack linear sequence similarity, making it difficult to leverage studies of one lncRNA to inform the understanding of others. We describe a k-mer-based hidden Markov model, hmSEEKR, that enables the scanning of transcriptomes for regions of non-linear sequence similarity to a query domain, without prior knowledge of where within the transcriptome the similarities may be located. When individual lncRNA domains were used as search features, hmSEEKR successfully identified regions in other RNAs that harbor non-linear sequence similarity and bind similar sets of proteins. Applying hmSEEKR to transcriptome-wide searches, we found that certain domains within the lncRNAs XIST, NEAT1, and MALAT1 exhibited widespread regional similarity to both lncRNA and protein-coding genes, while others were more unique, exhibiting similarity to ~100 genes or fewer. Combinatorial searches uncovered RNAs containing sequential matches to core functional domains of XIST and NEAT1, and eCLIP-inferred protein-interaction networks within these RNAs more closely resembled those of XIST and NEAT1, respectively, than would be expected by chance, suggesting the searches recovered RNAs with similar biological properties. Finally, within annotated sets of cis-activating and cis-repressive lncRNAs, we observed opposing enrichments for similarity to domains associated with transcription-promoting complexes and heterogeneous nuclear ribonucleoprotein (hnRNP) binding, respectively, suggesting the enriched sequences may contribute to regulatory functions. hmSEEKR can be applied with minimal training data and enables the a priori discovery of RNA domains that share nonlinear similarity, offering a sequence-informed approach to discover functional elements within noncoding transcriptomes.

13
Trinucleotide Distribution, Symmetry Elements and Formulation of Mirror Symmetry Index for G4 Motifs

Arya, A.; Datta, B.

2026-07-05 bioinformatics 10.64898/2026.07.05.736592 medRxiv
Top 0.1%
12.7%
Show abstract

Symmetry elements in nucleic acids are most strongly correlated with sites of biological function; however, their relevance to non-canonical structures remains underexplored. In this study, we demonstrate the presence and significance of trinucleotide symmetry elements within G-quadruplex (G4) motifs. Our central hypothesis is that the intra-strand mirror symmetry of trinucleotides has been evolutionarily selected to facilitate G4 formation builds on the established sequence-structure association of G-quadruplexes and the natural symmetry law governing nucleotide insertion during genome evolution. Using a conserved G4 motif in the first exon of the MTOR gene as a model, we showed remarkable trinucleotide symmetry preservation across primates and broader mammals, with functional G4 regions displaying locally elevated symmetry relative to the codon-biased exonic background. Analysis of experimentally validated oncogenic G4s, including c-MYC, BCL2, VEGF, and KRAS, revealed that mirror and reverse complement symmetries converge around biologically important G4s. To quantify this feature, we formulated two complementary descriptors: the mirror symmetry index (MSI) and its non-palindromic variant (nMSI). Across 14 oncogene-promoter wild-type G4s, the majority scored MSI [≥] 0.80 (mean 0.884), with only the loop-rich ATG7, BCR, and MDM2 motifs falling below this value, and the KRAS promoter G4 reached individual significance against its mononucleotide-preserving null distribution (p = 0.042). Most decisively, each wild-type G4 scored higher on MSI than its experimentally confirmed G4-abolished mutant in 12 of 14 paired comparisons (sign test, p = 0.0065; mean {Delta}MSI = +0.089, mean {Delta}nMSI = +0.192); the two reversals (BCL2 and HIF-1) are attributable to scrambled mutant controls that introduce more balanced trinucleotide compositions rather than to failure of the index. The directional trend was reproduced across three independently published datasets, with nMSI [≥] 0.50 separating G4-forming from non-G4 sequences at 77.8% sensitivity and 100% specificity, although the collective per-sequence signal from mononucleotide-preserving shuffles remained a non-significant trend (Stouffer combined Z = 1.197, p = 0.116). This first report of trinucleotide symmetry in G4 motifs posits that coordinated nucleotide insertion and quadruplet maintenance act as an evolutionary forcing mechanism that pre-organizes single strands for G4 folding.

14
HeartBioPortal 3.0: an integrated cardiovascular genomics knowledge environment for molecular, clinical and population-scale interpretation

Vand, K.; Badia, N.; Khomtchouk, B.; Janga, S. C.

2026-07-01 cardiovascular medicine 10.64898/2026.06.28.26356792 medRxiv
Top 0.1%
12.7%
Show abstract

Cardiovascular genomics is producing rapidly expanding genetic, molecular, phenotypic and clinical data, yet relevant evidence remains fragmented across resources and difficult to translate into actionable biological and ultimately translational knowledge. HeartBioPortal (HBP) is a browser-based cardiovascular knowledge environment that was developed to address this problem by organizing omics, variant, phenotype and clinical evidence centered around gene queries. Here we describe HBP 3.0, a major update that expands both the data architecture and interpretive interface. This update introduces DataHub, a reproducible data-engineering layer for source ingestion, standardization, variant-centered aggregation, provenance tracking and compact serving artifacts. The release integrates cardiovascular clinical practice guideline context through a graph-backed clinical knowledge layer; incorporates cardiovascular summary statistics from the Million Veteran Program and public aggregate resources; expands source-preserving population frequency, variant annotation and structural-variant; and adds gene profile, drug-discovery and protein-context layers. HBP 3.0 incorporates 594.3 million allele-frequency observations across 18.1 million rsIDs, 3.04 million exon-enriched structural-variant records, 66.9 thousand protein isoforms with 3.26 million non-exon protein feature annotations, 17,128 gene-drug records, and a clinical guideline knowledge graph with 42,895 entities and 106,304 relationships. The redesigned gene dossier view combines phenotype filtering, annotation composition, persistent selected-detail panels and exportable chart data in one workflow. HBP 3.0 is designed to help cardiovascular and eventually cardiometabolic researchers move from a genetic or genomic signal to biological knowledge and potentially clinical and therapeutic context while preserving source provenance and interpretive boundaries. Database URL: https://www.heartbioportal.com/

15
KozakExplorer: an interactive framework for genome-wide Kozak sequence analysis

Cokelaer, T.; Santi, A. M. M.; Pipoli da Fonseca, J.; Spaeth, G. F.

2026-06-26 bioinformatics 10.64898/2026.06.25.734688 medRxiv
Top 0.1%
12.5%
Show abstract

Translation initiation signals shape gene expression across all domains of life. In eukaryotes, nucleotide constraints surrounding the start codon are commonly described by the Kozak Consensus Sequence (KCS), whereas in bacteria and archaea, initiation frequently involves Shine--Dalgarno ribosome-binding motifs. Although these signals have been extensively characterized in model organisms, their large-scale diversity and evolutionary distribution remain incompletely explored. We present KozakExplorer, a reproducible framework for quantitative and comparative analysis of translation initiation contexts from genome assemblies and annotations. The software performs strand-aware extraction of start codon environments from FASTA and GFF3 files and applies information-theoretic metrics---including Kullback--Leibler (KL) divergence and information content (IC)---to measure positional nucleotide constraints relative to a background model. Derived summary statistics (Kozak Strength Index [KSI], maximum information content, peak position) convert motif patterns into interpretable per-genome signatures suitable for cross-species comparison. Our primary analysis covers 2,282 eukaryotic reference genomes, producing a standardized dataset of translation initiation metrics. Dimensionality reduction via t-SNE on per-position KL divergence, information content, and motif nucleotide frequencies reveals a structured eukaryotic KCS landscape with kingdom-level clustering and continuous variation in signal strength. A dedicated case study of 216 Apicomplexa genomes shows genus-level structure consistent with host range and phylogeny. An extended analysis across 25,344 reference genomes (22,253 bacteria, 809 archaea) places eukaryotic patterns in a global comparative framework, revealing transitions between sharply localized Kozak motifs and distributed Shine--Dalgarno-type signatures. Implemented within the open-source Sequana ecosystem, KozakExplorer is distributed as a Python module and an interactive web application that accepts local annotated assemblies, GenBank records, or NCBI RefSeq accessions, and exports all computed metrics, embeddings, and coordinates for downstream comparative and evolutionary genomics.

16
pysigscore: gene signatures scoring across bulk and single-cell transcriptomics

Giacomello, T.; Mazzara, S.; Abbruzzese, G.; Barberis, A.; tangherloni, a.; Buffa, F. M.

2026-08-09 bioinformatics 10.64898/2026.08.04.742537 medRxiv
Top 0.1%
12.5%
Show abstract

SummaryHigh-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures. Availability and ImplementationSource code is available at https://github.com/bioinformatics-hub/pysigscore. Contact: tommaso.giacomello@phd.unibocconi.it, francesca.buffa@unibocconi.it Supplementary informationSupplementary data are available at Bioinformatics online.

17
RD-OMICS: An Integrative Multi-Omics Data Inventory in Rare Diseases

Sun, S.; Wang, H.; Mathe, E. A.; Zhu, Q.

2026-07-03 bioinformatics 10.64898/2026.06.29.735296 medRxiv
Top 0.1%
12.4%
Show abstract

Rare diseases (RD) impact over 30 million individuals in the United States, yet fewer than 5% of the identified conditions have FDA-approved treatments. Progress in RD research is hindered by small patient cohorts, biological heterogeneity, and the fragmented, inconsistently annotated publicly available omics data, which limits integrative analysis and translational discovery. Here, we present RD-OMICS, a data inventory with integrated and structured RD omics data from Gene Expression Omnibus (GEO), in the form of a knowledge graph. We developed a metadata harmonization pipeline that combines rule-based mapping and large language model (LLM)-assisted semantic categorization. The graph-based data model was defined to integrate different types of data including disease conditions, experiments, samples, platforms, projects, and publications into a centralized inventory graph. In this preliminary study, 11,049 GEO series for 126 rare diseases were processed and integrated into RD-OMICS, which includes 375,930 individual biospecimen samples, 1,578 sequencing and array platforms, 10,938 biological projects. Case studies demonstrate the use of RD-OMICS in supporting rare disease research, omics cohort construction, and transcriptome-based drug repurposing for amyotrophic lateral sclerosis (ALS). RD-OMICS provides a scalable foundation for transforming fragmented omics data into a structured, harmonized and interoperable resource, facilitating therapeutic development and other translational discoveries in rare diseases.

18
Novel biologically relevant small RNA-sequencing alignment tool LevenMap for alignment to database of non-coding RNAs

Dlugas, H.; Dyson, G.; Dombkowski, A.; Kim, Y.; Gurdziel, K.; Boerner, J. L.; Bock, C.

2026-08-21 bioinformatics 10.64898/2026.08.14.742100 medRxiv
Top 0.1%
12.4%
Show abstract

A crucial aspect of the bioinformatics workflow in small RNA-sequencing is the alignment of reads to a database of reference ncRNAs. Alignment algorithms such as Bowtie, Burrows-Wheeler Aligner (BWA), and Spliced Transcripts Alignment to a Reference (STAR) - which are designed for aligning reads to a reference genome - are typically used. Aligning short RNA-sequenced reads to a database of non-coding RNAs (ncRNAs) is fundamentally a different task than aligning longer reads to a genome due to ncRNAs (i) having roughly the same number of nucleotides as the reads being aligned and (ii) being subsequences of other ncRNAs. To account for these differences, we developed the novel alignment algorithm LevenMap. Of all reads which exactly matched a reference ncRNA in a publicly available dataset, LevenMap aligned 100.0% of them to their respective ncRNA while all other aligners mapped less than 40% of these reads to their corresponding ncRNA. Furthermore, the mean ratio (length of read) / (length of corresponding reference ncRNA) of all aligned reads was 1.0 and 0.998 for LevenMap with at most zero and one mismatch(es) allowed, respectively; this ratio was no more than 0.51 for all other aligners. Overall, LevenMap is designed to account for the nuances of aligning small RNA-sequencing data to a database of reference ncRNAs and yields more biologically relevant counts compared to traditional aligners in this context. LevenMap is free and publicly available on GitHub: https://github.com/hdlugas/LevenMap.

19
Hidden sampling biases inflate performance in gene regulatory network inference

Stock, M.; Ratajczak, F.; Bertin, P.; Hoermanseder, E.; Bengio, Y.; Hartford, J.; Falter-Braun, P.; Heinig, M.; Tong, A.; Scialdone, A.

2026-07-14 bioinformatics 10.64898/2025.12.19.695616 medRxiv
Top 0.1%
12.3%
Show abstract

Accurate reconstruction of gene regulatory networks (GRNs) from single-cell transcriptomic data remains a major methodological challenge. Recent machine learning approaches, particularly graph neural networks and graph autoencoders, have reported improved performance, yet these gains do not consistently translate to realistic biological settings. Here, we show that a key reason for that is the way negative regulatory interactions are sampled for supervised training and evaluation. We find that widely used sampling strategies introduce node-degree biases that allow models to exploit trivial graph-structural cues rather than biological signals. Across multiple benchmarks, simple degree-based heuristics match or exceed state-of-the-art graph neural network models under these biased evaluation protocols. We further introduce a degree-aware sampling approach that eliminates these artifacts and provides more reliable assessments of GRN inference methods. Our results call for standardized, bias-aware benchmarking practices to ensure meaningful progress in supervised GRN inference from single-cell RNA-seq data.

20
CoTRA: an integrated R/Shiny framework for transparent bulk and single-cell RNA-seq analysis

Seemab, U.; Vainionpaa, K.; Tanoli, Z.; Leinonen, H. O.

2026-08-28 bioinformatics 10.64898/2026.08.25.747017 medRxiv
Top 0.1%
12.0%
Show abstract

Bulk and single-cell RNA sequencing (scRNA-seq) have become essential for investigating disease mechanisms and identifying diagnostic biomarkers. However, the growing volume of transcriptomic data remains difficult to reuse efficiently for many researchers. Downstream analysis often requires multiple statistical, visualization, and reporting tools, creating fragmented workflows that reduce transparency and reproducibility, particularly when analyzing scRNA-seq data. To address this gap, we developed CoTRA (Comprehensive Toolbox for RNA-seq Analysis), an open-source R/Shiny package for bulk and scRNA analysis. CoTRA integrates established methods into modular workflows, exposes parameters, and offers alternatives at selected stages. It supports bulk RNA-seq quality assessment, differential expression, annotation, enrichment, and reporting, as well as scRNA quality control, dimensionality reduction, clustering, marker identification, cell-type annotation, differential abundance, trajectory inference, pathway activity, and cell-cell communication. CoTRA runs on workstations or HPC environments without mandatory external data submission and was tested on Linux, Windows, and macOS. Compared with 14 other platforms for bulk RNA-seq/scRNA-seq, CoTRA supported 46 of 49 predefined functionality criteria. Tool validation using published rd10 retinal bulk RNA-seq identified 1,947 shared differentially expressed genes with concordant direction and strong log2 fold-change agreement. A retinal scRNA-seq case study demonstrated appropriate clustering, cell-type resolved analysis, and pathway activity scoring. CoTRA provides a graphical environment for bulk and single-cell RNA-seq analysis while retaining parameter transparency, methodological flexibility, and reproducible outputs. Strong concordance with the published bulk RNA-seq analysis supports the workflow consistency, while the single-cell case study demonstrates its applicability to advanced scRNA-seq analysis. The source code is freely available at https://github.com/UmairSeemab/CoTRA.